Skip to content

ci: clean up recurring nightly optional-deps failures - #2968

Merged
leofang merged 10 commits into
NVIDIA:mainfrom
leofang:fix-nightly-optional-deps-2748
Sep 30, 2026
Merged

leofang merged 10 commits into
NVIDIA:mainfrom
leofang:fix-nightly-optional-deps-2748

Conversation

@leofang

@leofang leofang commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Cleanup pass for the recurring failures tracked in #2748, grouped by affected library.

numba-cuda

  • Pin numpy<2.5 in ci/tools/run-tests for nightly-numba-cuda. numba-cuda 0.30.4 calls np.row_stack, which NumPy 2.5 removed, so every nightly numba-cuda job crashed at collection. Tracked upstream in NVIDIA/numba-cuda#907.
  • Pin pytest<9 as a second pip install in ci/tools/run-tests. numba-cuda's subTest usage breaks under pytest 9 + xdist (NVIDIA/numba-cuda#637). Split into a second install because cuda_core's test-cuXX group hard-pins pytest==9.1.0, which pip's resolver can't satisfy alongside pytest<9 in the same solve.
  • Switch the test runner from python -m numba.runtests to pytest, mirroring the nightly-numba-cuda-mlir step: run-tests exposes NUMBA_CUDA_VER, the workflow checks out NVIDIA/numba-cuda at v${NUMBA_CUDA_VER}, and the test step runs from numba-cuda-released/testing/ so upstream's testing/pytest.ini (consider_namespace_packages = true + --pyargs numba.cuda.tests) drives discovery. Avoids the namespace-package/pytest interaction that was pushing numba.runtests in the first place.
  • Set NUMBA_CUDA_TEST_WHEEL_ONLY=1 in wheels mode so tests requiring CTK executables self-skip. Same env var upstream numba-cuda uses in its own wheels-mode CI; matches the NUMBA_CUDA_MLIR_TEST_WHEEL_ONLY gate already in the mlir step.
  • -k "not TestIpc" — numba-cuda 0.30.4 accesses CUipcMemHandle.reserved, which cuda-bindings intentionally removed; numba-cuda is EOL upstream so no fix is coming. Only fires on linux-64 x86_64 — TestIpcMemory and TestIpcStaged are stacked with @linux_only/@skip_on_arm/@skip_on_wsl2 and auto-skip elsewhere.

numba-cuda-mlir

  • Add the cccl extra to cuda-toolkit in ci/tools/run-tests for nightly-numba-cuda-mlir. NVRTC was failing with catastrophic error: cannot open source file "nv/target" while compiling cooperative_groups/details/info.h and curand_kernel.h; the cccl extra pulls nvidia-cuda-cccl, which drops the missing header at nvidia/cu13/include/nv/target.
  • --deselect TestCudaDeviceRecordWithRecord::test_device_record_copy. The fixture uses np.recarray (uninitialized memory); when the float32 field happens to encode a NaN, np.testing.assert_equal fails on NaN != NaN even though the copy round-trip is correct. Filed upstream as NVIDIA/numba-cuda-mlir#341.

cuda-core

  • Version-gated --deselect for test_frozen_driver_table_covers_all_curesult_members on released cuda-core <= 1.2.1. Main cuda-bindings 13.4.1 exposes 3 new CUresult members (CUDA_ERROR_MULTICAST_RESOURCE_FULL, CUDA_ERROR_INSUFFICIENT_LOADER_VERSION, CUDA_ERROR_FABRIC_NOT_READY) not covered by the frozen fallback table those releases ship. The test was already reverted from main by [no-ci] Refreeze enum tables (revert 2383, add comment) #2792, so this deselect drops automatically once the next cuda-core release ships that revert. Mirrors the existing v1.0.1 NvlinkVersion deselect pattern in the same step.

…mlir

Two targeted fixes for the recurring failures tracked in NVIDIA#2748:

- nightly-numba-cuda: pin numpy<2.5. numba-cuda 0.30.4 references
  np.row_stack, removed in NumPy 2.5, so every nightly numba-cuda
  job crashes at collection. Tracked upstream in
  NVIDIA/numba-cuda#907; drop the cap once a fixed wheel is on
  PyPI.

- nightly-numba-cuda-mlir: add cccl to cuda-toolkit extras. NVRTC
  currently fails with "catastrophic error: cannot open source
  file 'nv/target'" when compiling cooperative_groups/details/info.h
  and curand_kernel.h; the cccl extra pulls nvidia-cuda-cccl, which
  drops the missing header at nvidia/cu13/include/nv/target.
@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the CI/CD CI/CD infrastructure label Sep 30, 2026
@leofang leofang changed the title ci: pin numpy&lt;2.5 for nightly-numba-cuda and add cccl extra for numba-cuda-mlir ci: pin numpy<2.5 for nightly-numba-cuda and add cccl extra for numba-cuda-mlir Sep 30, 2026
test_frozen_driver_table_covers_all_curesult_members was removed on
main by NVIDIA#2792 (2026-09-09) but neither cuda-core-v1.2.0 nor v1.2.1
carry the revert. main cuda-bindings 13.4.1 adds three new CUresult
members (MULTICAST_RESOURCE_FULL, INSUFFICIENT_LOADER_VERSION,
FABRIC_NOT_READY) not present in the frozen table those releases
ship, so the nightly-cuda-core job trips it on every run.

Mirror the existing v1.0.1 NvlinkVersion pattern with a
version-gated --deselect that drops automatically once the next
cuda-core release ships the NVIDIA#2792 revert.
DESELECTS+=(
--deselect 'tests/test_utils_enum_explanations_helpers.py::test_frozen_driver_table_covers_all_curesult_members'
)
fi

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unfortunately #2792 did not catch any cuda.core v1.2.x release and only lives on the main branch, so we have to fix the nightly CI this way.

@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

/ok to test feb41f7

@leofang leofang self-assigned this Sep 30, 2026
@leofang leofang added this to the cuda.core 1.3.0 milestone Sep 30, 2026
@leofang leofang added the bug Something isn't working label Sep 30, 2026
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor
Doc Preview CI
Preview removed because the pull request was closed or merged.

nightly-numba-cuda: switch from `python -m numba.runtests` to `pytest`
(numba-cuda's own CI uses pytest; PR NVIDIA#1987 landed the runtests form as a
bring-up fix) so we can skip individual tests. Two workarounds:

- pin pytest<9 (matches upstream's numba-cuda cap; NVIDIA/numba-cuda#637
  covers the subTest breakage on 9).
- --ignore-glob "**/test_ipc.py" — numba-cuda v0.30.4 accesses
  CUipcMemHandle.reserved, which cuda-bindings intentionally removed;
  numba-cuda is EOL upstream so no fix is coming (NVIDIA#2748).

nightly-numba-cuda-mlir: add --deselect for
TestCudaDeviceRecordWithRecord.test_device_record_copy — the WithRecord
fixture uses np.recarray (uninitialized memory), so when the float32
field encodes a NaN, `np.testing.assert_equal` trips on NaN != NaN even
though the copy round-trip is correct. Filed upstream as
NVIDIA/numba-cuda-mlir#341.
Keeps pip constraints in the install step alongside numpy<2.5,
instead of a separate `pip install` in the workflow.
@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

/ok to test f0682bf

@leofang leofang added the P0 High priority - Must do! label Sep 30, 2026
cuda_core's test-cuXX group hard-pins pytest==9.1.0, so pip can't
satisfy that and pytest<9 in one solve. Split the numba-cuda pytest
pin into a second `pip install "pytest<9"` after the main install.
@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

/ok to test 31f7af7

Default `prepend` import mode walks up until it finds a directory
without `__init__.py`. numba-cuda's on-disk layout is
`site-packages/numba_cuda/numba/cuda/tests/...`, and `numba/` here is
a namespace subdir with no `__init__.py`, so pytest treats it as the
rootpath and imports each test file as `cuda.tests.<...>`. That
collides with cuda-bindings' top-level `cuda` package (which has no
`tests` subpackage), giving `ModuleNotFoundError: No module named
'cuda.tests'` at collection.

`--import-mode=importlib` sidesteps the sys.path walk and imports
each test module by its correct dotted path.
Same shape as the numba-cuda-mlir step. run-tests exposes
NUMBA_CUDA_VER; the workflow checks out NVIDIA/numba-cuda at the
matching tag and runs pytest from numba-cuda-released/testing/, so
upstream's testing/pytest.ini kicks in
(consider_namespace_packages=true + --pyargs numba.cuda.tests) and
we stop caring about the site-packages namespace-package layout.

Also mirror the wheel-only gate: in wheels mode, set
NUMBA_CUDA_TEST_WHEEL_ONLY=1 so numba-cuda skips tests that need
CTK executables. Binary-generated tests self-skip via
`@unittest.skipIf(not NUMBA_CUDA_TEST_BIN_DIR, ...)`, so no
`make -j` step is needed — mirrors mlir, which has no Makefile.

Drops the --import-mode=importlib workaround from the prior commit;
pytest.ini's consider_namespace_packages does the job.
@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

/ok to test 3bc4acc

pytest's --ignore-glob="**/test_ipc.py" didn't match through the
--pyargs-resolved site-packages path (5 tests still ran and failed
with the CUipcMemHandle.reserved AttributeError on linux-64). Match
by class name instead — the affected tests live in TestIpcMemory
and TestIpcStaged, both starting with "TestIpc", so this is
unambiguous.
@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

/ok to test b53d482

@leofang leofang changed the title ci: pin numpy<2.5 for nightly-numba-cuda and add cccl extra for numba-cuda-mlir ci: clean up recurring nightly optional-deps failures Sep 30, 2026

@brandon-b-miller brandon-b-miller left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

numba-cuda/numba-cuda-mlir related changes LGTM.

Comment thread .github/workflows/test-wheel-linux.yml
@leofang

leofang commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Since the CI was green, let me admin-merge to save resource.

@leofang
leofang marked this pull request as ready for review September 30, 2026 14:30
@leofang
leofang merged commit 4b9dcbd into NVIDIA:main Sep 30, 2026
5 checks passed
@leofang
leofang deleted the fix-nightly-optional-deps-2748 branch September 30, 2026 14:30
github-actions Bot pushed a commit that referenced this pull request Oct 1, 2026
Removed preview folders for the following PRs:
- PR #2939
- PR #2953
- PR #2962
- PR #2966
- PR #2968
- PR #2969
- PR #2972
- PR #2974
- PR #2976
- PR #2977
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working CI/CD CI/CD infrastructure P0 High priority - Must do!

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants